Papers with Wikipedia edits

3 papers
Generating Inflectional Errors for Grammatical Error Correction in Hindi (2020.aacl-srw)

Copied to clipboard

Challenge: Automated grammatical error correction is a data-heavy task . indic languages have a relatively low amount of digitized content and complex morphology .
Approach: They generate a corpus of inflectional errors for training neural networks to correct grammatical errors in Hindi.
Outcome: The proposed model trains on a corpus of inflectional errors extracted from Wikipedia edits.
WIKIBIAS: Detecting Multi-Span Subjective Biases in Language (2021.findings-emnlp)

Copied to clipboard

Challenge: a particular type of bias is subjective bias, which introduces improper attitudes or presents a statement with the presupposition of truth.
Approach: They propose to annotate a Wikipedia edits corpus with 4,000 sentence pairs to detect subjective bias.
Outcome: The proposed dataset can be used as a research benchmark and generalize to multiple domains.
The Million Authors Corpus: A Cross-Lingual and Cross-Domain Wikipedia Dataset for Authorship Verification (2025.findings-acl)

Copied to clipboard

Challenge: Authorship verification (AV) is a crucial task for identity verification, accountlinking, historical linguistics, and AI-generated text identification.
Approach: They propose to use Wikipedia's Million Authors Corpus to examine authorship verification models on a broad scale.
Outcome: The proposed dataset includes 60.08M textual chunks, contributed by 1.29M Wikipedia authors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations